Papers with speech dataset
Crowdsourcing Speech Data for Low-Resource Languages from Low-Income Workers (2020.lrec-1)
Copied to clipboard
Basil Abraham, Danish Goel, Divya Siddarth, Kalika Bali, Manu Chopra, Monojit Choudhury, Pratik Joshi, Preethi Jyoti, Sunayana Sitaram, Vivek Seshadri
| Challenge: | Existing platforms collect labelled speech data from urban speakers whose dialects are often very different from low-income users. |
| Approach: | They propose to collect labelled speech data directly from low-income workers . they collect 109 hours of data from 36 participants in the Marathi language . |
| Outcome: | The proposed approach can provide valuable supplemental earning opportunities to low-income rural and urban workers. |
TV-AfD: An Imperative-Annotated Corpus from The Big Bang Theory and Wikipedia’s Articles for Deletion Discussions (2020.lrec-1)
Copied to clipboard
| Challenge: | Detecting imperatives in oral and written communication is difficult when the user doesn't use the expected forms. |
| Approach: | They created an imperative corpus with dialogues from The Big Bang Theory and Wikipedia comments from Wikipedia . they manually annotated imperatives and used a syntax-based classifier to extract 10,624 statements that may be imperative. |
| Outcome: | The proposed model performs better in the written data compared to speech data, but has a low precision and recall for speech data. |